Papers with low-resource language scenarios

2 papers
A Semantic Uncertainty Sampling Strategy for Back-Translation in Low-Resources Neural Machine Translation (2025.acl-srw)

Copied to clipboard

Challenge: Back-translation methods rely on large-scale parallel corpora to enhance performance, but ignore the semantic quality of monolingual data.
Approach: They propose a method which prioritizes sentences with higher semantic uncertainty as training samples by computationally evaluating the complexity of unannotated monolingual data.
Outcome: The proposed method improves translation accuracy and fluency by +1.7 on all three translation tasks.
A Multilingual Topic Model for Learning Weighted Topic Links Across Corpora with Low Comparability (D19-1)

Copied to clipboard

Challenge: Existing models implicitly assume that documents in different languages are highly comparable, a false assumption.
Approach: They propose a multilingual topic model that learns weighted topic links and connects cross-lingual topics only when the dominant words defining them are similar.
Outcome: The proposed model outperforms existing models in low-resource language tasks and outperformed LDA and previous models in classification tasks using documents’ topic posteriors as features.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations